Back

JAMIA Open

Oxford University Press (OUP)

Preprints posted in the last 30 days, ranked by how well they match JAMIA Open's content profile, based on 42 papers previously published here. The average preprint has a 0.07% match score for this journal, so anything above that is already an above-average fit.

1
Performance, Generalizability, and Fairness of a Peripheral Artery Disease Detection Model Across Patient Phenotypes and Health Systems

Kallis, K.; Quitevis, C. R.; Ramsis, M.; Kabutey, N.-K.; Conte, M. S.; Rowe, V. L.; Humphries, M. D.; Hernandez-Boussard, T.; B. Malas, M.; Ross, E. G.

2026-08-22 health informatics 10.64898/2026.08.19.26360861 medRxiv
Top 0.1%
12.6%
Show abstract

Background Peripheral artery disease (PAD) is a major cause of cardiovascular events but remains underdiagnosed. Electronic health record (EHR)-based machine learning models show promise for earlier detection, but developing generalizable and fair models across diverse populations remains challenging. Methods Using the University of California Health Data Warehouse, containing EHR data from five health systems, we identified patients with and without PAD. We used unsupervised clustering to define PAD phenotypes and trained a LightGBM classifier using 14,023 features spanning demographics, comorbidities, medications, laboratory values, healthcare utilization, and diagnosis, procedure, and medication codes. We evaluated performance overall and across demographic groups and phenotypes, and assessed fairness using selection rates and subgroup differences in true- and false-positive rates. Results The study included 33,739 cases and 33,739 matched controls. Clustering identified four phenotypes: patients with limited healthcare documentation (cluster 1), younger patients with severe metabolic disease (cluster 2), patients with a traditional atherosclerotic risk profile (cluster 3), and frail elderly patients with multimorbidity (cluster 4). Overall, the model demonstrated consistent performance across institutions (AUROC 0.76?0.79; AUC-PR 0.76?0.79) with well-calibrated probabilities. Performance was similar across genders, with modest variation by race and age, and was stronger in clusters 2?4. Cluster 2 demonstrated the highest sensitivity (TPR 0.87, 95% CI 0.87?0.88), while cluster 1 showed the lowest performance (TPR 0.40, 95% CI 0.39?0.41). Conclusions The EHR-based PAD detection model demonstrated consistent performance across five health systems. Phenotypic clustering revealed clinically meaningful differences in model performance adding an additional consideration in ML fairness and performance evaluations.

2
DBToken: A Database Tokenizer for Medical Event Foundation Models

Shin, I.; McCann, K.; Marino, G.; Siam, U. T.; Li, H.; Stutz, E.; Edara, R.; Loza, A. J.

2026-08-21 health informatics 10.64898/2026.08.18.26360487 medRxiv
Top 0.1%
12.5%
Show abstract

Objectives Transformer models for electronic health records require converting clinical data into token sequences, however standardized tokenization and evaluation frameworks are lacking. We introduce DBToken, an open-source library, and bits-per-row (BPR), a metric for comparing tokenization strategies. Materials and Methods DBToken accepts Medical Event Data Standard (MEDS)-compatible input and supports multiple text, numeric, and temporal tokenization strategies. BPR extends the bits-per-byte metric used in language models to enable comparison across tokenization strategies. Results DBToken efficiently tokenized data across configurations. BPR identified the vocabulary size associated with the best clinical outcome performance and localized differences in numeric tokenization performance by token class. Discussion Optimal tokenization strategies for medical foundation models are a subject of active research. DBToken enables reproducible tokenization experiments, while BPR efficiently screens vocabulary sizes and numeric representations before downstream evaluation. Conclusion DBToken and the BPR metric provide open-source infrastructure for reproducible EHR tokenization and cross-strategy evaluation.

3
Large Language Models Generate Stigmatizing Language During Reasoning Over Real-World Clinical Data

Yang, Y.; Gu, B.; Hathaway, D. B.; Wyss, R.; Marengo, L.; Gibbons, J. B.; Lyndon, S.; Wu, J.; Chen, Q.; Liu, N.; Wang, P. S.; Celi, L. A.; Bates, D. W.; Lin, J.; Zhou, L.; Yang, J.

2026-08-14 health informatics 10.64898/2026.08.12.26360210 medRxiv
Top 0.1%
10.6%
Show abstract

Stigmatizing language in clinical documentation, which conveys negative stereotypes, attitudes, or judgments toward patients, is a recognized source of documentation bias and is associated with poorer care and adverse health outcomes. Although prior stigma-related research has focused on clinician-written EHR notes, the increasing use of large language model (LLM)-generated documentation in clinical workflows raises new concerns about its potential to reproduce or amplify bias and affect patient safety. In this study, we conducted a large-scale assessment of stigmatizing language in LLM-generated reasoning text on 35 real-world clinical tasks across 107 LLMs. We applied a psychiatrist-validated, natural language processing (NLP) system to detect stigma terms in LLM reasoning text and quantified stigma rates of LLM-generated reasoning texts across 3,745 model-task pairs. Results showed that stigma rates ranged from 0% to 33.33%, with 84.06% of pairs containing stigma terms. Open-source models and reasoning models showed statistically higher stigma rates than proprietary (1.97% vs. 1.60%; p < 0.01) and non-reasoning models (2.35% vs. 1.70%; p < 0.0001), while the stigma rate difference between the general and medical models is not statistically significant (2.00% vs. 1.80%; p = 0.26). Stigma rates of LLM outputs correlated negatively with task accuracy (r = -0.304; p < 0.001) and positively with input clinical-text stigma (r = 0.569; p < 0.001), with 19.76% of model-task pairs amplifying stigma in the original input notes. Applying prompt engineering as a destigmatizing approach helped reduce model stigma rates by as much as 91.91% without affecting the model performance. This study shows that stigmatizing language generation is common but reducible during LLMs' reasoning traces, suggesting that well-implemented approaches for LLM monitoring and destigmatizing will be essential for healthcare systems to implement.

4
Machine Learning-Based Prediction of Maternal Morbidity across Heterogeneous Populations in the United States using Sequential Modeling of the All of Us Dataset

Zhuang, H.; Zakama, A.; Heller, K.; Faulkner, S.; Gollub, B.; Young-Lin, N.; Chen, I. Y.; Asiedu, M.

2026-08-31 obstetrics and gynecology 10.64898/2026.08.25.26360552 medRxiv
Top 0.2%
8.0%
Show abstract

In this work, we demonstrate the unprecedented value of NIH's "All of Us Research Program" (AoURP) dataset in studying maternal morbidity and building predictive machine learning (ML) models across heterogeneous populations in the United States. We developed robust and data-driven preprocessing pipelines to curate a longitudinal, multi-site, multimodal, and demographically diverse pregnancy dataset (20,253 subjects; 27,525 pregnancy episodes) from AoURP data, using electronic health records (EHR) (Conditions, Labs, Measurements) and survey responses (Social Determinant of Health (SDoH)), focusing on 7 crucial maternal health adverse outcomes. After characterizing data quality, missingness, and heterogeneity, we performed statistical correlation analysis to identify risk factors. We subsequently developed XGBoost and sequential LSTM models to predict the adverse outcomes, reaching state-of-the-art performance for multiple outcomes. We conducted model interpretability post-hoc analysis to understand success points and fairness analysis to evaluate implications for socio-economic disparities. Four practicing physicians reviewed the set of statistically significant and ML model identified features to assess their clinical validity and novelty. Most features identified through either statistical correlations or ML feature importance analysis aligned with known clinical risk factors. Several features were identified that the ML models used but that are not currently used in clinical practice and may merit further clinical investigation. Fairness analysis revealed certain associations with SDoH and age highlight areas that warrant continued monitoring. Overall, we demonstrate that meaningful populational level patterns can be extracted, and high-performing machine learning models can be trained on this longitudinal, diverse, multi-site dataset. Important risk features, particularly novel ones identified, if validated, could inform new strategies for maternal care or enable development and validation of outcome-specific, clinically deployable ML models.

5
A Human-in-the-Loop Large Language Model System Based on the Model Context Protocol for Differential Diagnosis from Electronic Medical Records and Literature

Lim, H.; Yi, H.; Yoon, J. Y.; Kwon, H.; Lee, D.; Kim, N.

2026-08-21 health informatics 10.64898/2026.08.18.26359085 medRxiv
Top 0.2%
7.8%
Show abstract

Diagnostic errors, including misdiagnoses and delayed clinical diagnoses, could affect outcomes of a significant patient population, particularly individuals presenting with rare diseases or non-specific symptoms. From rule-based diagnostic decision supporting systems (DDSS) to large language model (LLM) based tools for clinical reasoning have been developed to address these limitations. However, existing DDSS are often proprietary and difficult to integrate, and recent LLM-based tools remain hindered by operational challenges such as cost, resources constraint, and privacy concerns. Moreover, existing systems interpret electronic medical records (EMR) and generate diagnoses separately, limiting continuous evidence-based analysis and imposing repeated clinician involvement. In this paper, we present DDx-Finder, an open-source framework that leverages Model Context Protocol (MCP) servers for direct EMR and literature access, enabling prompt-driven clinical state extraction and reliable case-report re- trieval via generating searching query by LLM, while addressing limitations related to resource demands and privacy concerns. A clinical case study demonstrates the systems feasibility and its potential to provide accessible, transparent, and systematic differential diagnostic support for complex cases.

6
The EHR Density Index: A new method to control for EHR data inconsistency across patients

Bhatia, A.; Lash, S.; McIntee, T.; Pfaff, E.

2026-08-06 health informatics 10.64898/2026.08.03.26359595 medRxiv
Top 0.2%
7.6%
Show abstract

Electronic health record (EHR) data vary substantially in documentation density across patients, independent of disease burden. Existing tools such as the Charlson Comorbidity Index (CCI) and Elixhauser Comorbidity Index measure disease burden but do not capture differences in data volume, leaving a common source of bias unaddressed in EHR-based analyses. To address this gap, we developed the EHR Density Index (EDI), which characterizes the quantity, depth, and breadth of EHR data per patient per year, normalized by utilization patterns, using records from 24,987 adult patients at UNC Health (2018 - 2024). The EDI combines a utilization cluster assigned via Gaussian Mixture Model with within-cluster residuals quantifying documentation volume across four clinical domains. Four interpretable clusters emerged; while CCI predicted cluster membership, its associations with within-cluster residuals were weak, confirming the EDI captures dimensions of the patient record distinct from disease burden. The EDI is intended as a covariate to address documentation density as a source of confounding in real-world data-driven research.

7
Early Detection of Erythropoietic Protoporphyria Using Sequential Machine Learning on Longitudinal Electronic Health Records

Ayati, A.; Onal, G.; Sur, A.; Azzam, S.; Wang, B.; Rudrapatna, V. A.

2026-08-17 gastroenterology 10.64898/2026.08.15.26360514 medRxiv
Top 0.2%
7.6%
Show abstract

Objective: Erythropoietic protoporphyria (EPP) is a rare photodermatosis marked by multi-year diagnostic delays. We developed and externally validated machine learning models to identify patients with EPP earlier from longitudinal electronic health record (EHR) data and estimate undiagnosed disease burden. Materials and Methods: In a retrospective case-control study at two San Francisco health systems, an academic referral center (UCSF) and a safety-net hospital (ZSFG) we identified 74 confirmed EPP cases using combined diagnostic coding, biochemical criteria, and specialty chart review. Symptom-enriched controls were sampled at a 40:1 ratio. Longitudinal diagnoses, laboratory results, medications, procedures, and encounters preceding the outcome date were modeled with a gradient-boosting classifier (CatBoost) and a state-space sequence model (MAMBA). The best model was deployed across the UCSF population and externally validated at ZSFG without retraining. Results: On the UCSF held-out test set (n=1,865; 43 cases), MAMBA outperformed CatBoost (AUC ROC 0.91 vs 0.89; average precision 0.42 vs 0.27; precision 65% vs 20%), flagging cases a median of 229 days before documented diagnosis. Deployed across 297,967 symptom-compatible patients, it identified 310 high-risk individuals, implying a prevalence approaching genetic estimates. External validation at ZSFG showed attenuated performance (AUC ROC 0.72; average precision 0.10) while preserving early detection (median 264 days). Discussion: A sequence model integrating temporal EHR signals detected EPP months before clinical recognition, corroborating genetic evidence of substantial underdiagnosis. Cross-site attenuation reflects population and documentation differences and underscores the need for local recalibration. Conclusion: Longitudinal EHR-based machine learning can shorten EPP diagnostic delay and prioritize patients for confirmatory testing, supporting proactive rare-disease case finding.

8
Counterfactual Analysis of Executable Clinical Decision Logic

Maleki, C.; Bertrand, Y.; Gailly, F.

2026-08-07 health informatics 10.64898/2026.08.05.26359737 medRxiv
Top 0.2%
7.3%
Show abstract

Clinical recommendations are often expressed in narrative form, which limits their direct execution, auditability, and patient-specific interpretation. This paper presents a hybrid decision-support framework that combines Decision Model and Notation (DMN), survey-weighted rule-ensemble learning, and counterfactual sensitivity analysis. The framework is evaluated using an NHANES-derived fasting cohort for classification of documented diabetes status. The full fasting analysis cohort contained 2,582 participants, and a non-diagnostic laboratory subgroup, Gate0, contained 2,111 participants. On untouched test data, the rule-ensemble model achieved ROC-AUC and PR-AUC values of 0.959 and 0.873 in the full fasting cohort and 0.861 and 0.499 in Gate0. Four clinically interpretable candidate rules were selected using validation data only. A nonnegative survey-weighted logistic model removed one redundant rule and converted the remaining three binary activations into an auditable DMN score and model-estimated probability. The final DMN achieved ROC-AUC 0.769, PR-AUC 0.153, and Brier score 0.029 in the untouched Gate0 test set. In small rule-defined test subgroups, hypothetical five-unit BMI reductions lowered mean model-estimated probability by 2.40 to 5.89 percentage points when one or more BMI thresholds were crossed. These findings characterize policy sensitivity rather than causal effects and require external validation.

9
A Guided AI Framework for Customizable and Efficient Harmonisation to the OMOP Common Data Model

Nehra, N.; Swami, R.; Dadi, D.; Mishra, R.; Sharma, U.; Verma, P.; Sen, M.; Dhruw, N. K.; Jha, A. K.

2026-08-12 bioinformatics 10.64898/2026.08.07.742453 medRxiv
Top 0.2%
7.2%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWGetting clinical data from different sources to "talk" to each other within the OMOP Common Data Model (CDM) is arguably the most tedious part of multi-center research. While this integration is essential, the transformation process is frequently a manual grind, requiring a rare overlap of deep clinical knowledge and technical expertise. In this paper, we present a framework designed to alleviate some of the burden on the researcher by automating data harmonization through two distinct steps: structural schema mapping and terminological standardization. For the structural piece, we moved away from "black box" logic in favor of a stateful workflow managed by large language models (LLMs) and directed acyclic graphs. By profiling EHR data at the source, our system generates context-aware dictionaries that offer ranked mapping suggestions alongside confidence scores. While our benchmarking showed a 97.5% agreement rate at the schema level and an 84% agreement rate at the value level when compared with human experts, the system appears most effective when treated as a "co-pilot" rather than a total replacement for human oversight. To handle value-level standardization, we implemented a hybrid search strategy that pairs the semantic depth of SapBERT embeddings with the literal precision of fuzzy string matching. By using FAISS for rapid similarity retrieval, the engine attempts to resolve messy or "noisy" clinical descriptions to standard OMOP concepts. This approach seems particularly promising for handling the non-standardized labels that often plague smaller, local datasets. Ultimately, our results suggest that this guided approach can shift the timeline for OHDSI-compliant warehousing from weeks of manual curation to a more manageable and scalable pipeline, potentially lowering the barrier to entry for smaller research teams.

10
Global Adoption of openEHR Clinical Data Repositories: A Vendor and Community Survey

Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.

2026-08-31 health informatics 10.64898/2026.08.27.26361529 medRxiv
Top 0.2%
7.1%
Show abstract

The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.

11
Development and Internal Validation of a Large Language Model Pipeline for Multi-Label Classification of Patient Portal Messages

Steitz, B. D.; Ogunsan, O. O.; Ancker, J. S.; Carlson, B. R.; Gaynor, L. S.; Higashi, R. T.; Morrow, E. L.; Reese, T. J.; Romano, R. R.; Stern, S.; Turer, R. W.; Rosenbloom, S. T.; Wright, A.

2026-08-17 health informatics 10.64898/2026.08.14.26360460 medRxiv
Top 0.2%
6.7%
Show abstract

Objectives: Characterizing patient portal message content at scale can help target efforts to manage administrative work. We developed and validated a large language model (LLM) pipeline for multi-label classification of messages using an expert-derived topic taxonomy, then characterized topic distribution across a two-year corpus. Materials and Methods: We studied all medical advice request messages sent to ambulatory clinicians at an academic medical center from 2024-2025. We convened an expert panel that derived an 11-category taxonomy through a modified Delphi process. Two annotators labeled 750 randomly selected messages (Cohen kappa 0.80), holding out 500 for evaluation. The pipeline used GPT-4o-mini in a zero-shot prompt. On the held-out set, we measured micro- and macro-averaged precision, recall, and F1, and label stability across runs. We then characterized topic distribution and co-occurrence across the corpus. Results: The pipeline achieved micro- and macro-averaged F1 of 0.89 and 0.86. Labels were identical across runs for 93.6% of messages. Across 2.4 million messages, content concentrated on a few topics. The two most common topics, Problems & Management and Medications & Prescriptions, were present in 67.9% of messages, and the four most common in 93.9%. 51.7% of messages addressed multiple topics. Discussion and Conclusion: The pipeline classified patient message topics accurately and stably across millions of messages. Message content was concentrated within a small number of topics, highlighting opportunities for targeted interventions and enabling more efficient triage, routing, and patient-facing support.

12
Toward Transportable Acute Kidney Injury Prediction: An Explainable XGBoost Model with Temporal Validation Using MIMIC-IV

Okundaye, D. O.; Isiekwene, C. C.

2026-09-03 health informatics 10.64898/2026.09.01.26360393 medRxiv
Top 0.2%
6.7%
Show abstract

Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.

13
Explainable Clinician-Supervised Artificial Intelligence as an Implementation Framework for Cardiovascular-Kidney-Metabolic Population Health: Synthetic Data Validation of the CHAPERONE-CKM Framework

Vijay, A.; Govind, N.; Moorthy, A.; Dunn, P.; Lababidi, Z.; Jones, S.; Stahlberg, M.; Ibrahim, S.; Koochek, K.; Shah, K. S.; Schulhauser, R.; Lerma, E. V.; Nair, L.; Livi, J.; Kalra, D. K.; Wadwekar, D.; Gulllett, W.; Vijayaraghavan, K.

2026-08-19 health informatics 10.64898/2026.08.17.26360643 medRxiv
Top 0.2%
6.6%
Show abstract

Abstract Background: Cardiovascular-kidney-metabolic (CKM) syndrome is an increasingly prevalent multisystem condition associated with morbidity, fragmented care, recurrent hospitalization, and rising healthcare costs. While cardiovascular risk models estimate future disease risk, fewer frameworks support multidisciplinary CKM care, clinician decision-making, and population health management. Synthetic data environments can assess implementation readiness while preserving privacy. Methods: We validated the explainable, clinician-supervised CHAPERONE-CKM framework using a reproducible synthetic cohort of 10,090 simulated patients with 128 demographic, laboratory, imaging, treatment, and healthcare utilization variables across the CKM continuum. Synthetic data generation was separated from framework evaluation through probabilistic modeling and independent validation to reduce deterministic relationships. The framework generated CKM stage assignments, implementation priorities, clinician-readable rationales, multidisciplinary referral pathways, and guideline-directed therapy prompts. Evaluation focused on implementation readiness, consistency, calibration, subgroup stability, fairness, workflow simulation, and explainability. Results: The synthetic population represented CKM-related conditions including diabetes (52%), hypertension (65%), chronic kidney disease (20%), heart failure (32%), and prior CKM hospitalization (27%). The framework showed stable internal behavior across demographic and clinical subgroups, favorable calibration, and biologically plausible prioritization of advanced CKM disease. Workflow simulations suggested earlier identification of patients suitable for multidisciplinary review, therapy optimization, and coordinated care compared with reactive workflows. Traditional performance metrics supported framework behavior but were treated as secondary evidence rather than proof of clinical effectiveness. Conclusions: In a synthetic validation environment, the CHAPERONE-CKM framework demonstrated implementation readiness, transparent decision pathways, and compatibility with multidisciplinary CKM population health management. These findings are an early translational milestone, not clinical validation, and support external validation, prospective implementation studies, and Learning Health System integration to assess effects on care delivery, equity, and value-based outcomes.

14
Who seeks care, and what gets measured? Understanding the distinct mechanisms behind visit and observation processes in multi-center electronic health records

Yang, C.-H.; Salvatore, M.; Lu, H.; Zhu, Z.; Tennant, P.; Shi, X.; Ohno-Machado, L.; Khera, R.; Gross, C.; Li, F.; Mukherjee, B.

2026-08-14 health informatics 10.64898/2026.08.12.26360236 medRxiv
Top 0.3%
6.4%
Show abstract

Electronic health record (EHR)-linked cohorts support association, prediction, and causal studies using longitudinally measured markers of health. However, a lab biomarker measurement is recorded only when a patient first has a medical encounter (visit process) and, a clinician orders the corresponding test and the patient follows through (observation process). These two stages may induce informative presence (IP) and informative observation (IO), respectively. Yet their drivers remain largely uncharacterized, despite evidence that understanding this recording mechanism is essential for selecting appropriate strategies for downstream analysis that treat these markers as longitudinally measured outcomes. We characterize this two-stage recording hierarchy using a stochastic recurrent-event model for the outpatient visit process and a visit-process-weighted generalized estimating equation model for biomarker recording conditional on an outpatient visit. We characterize descriptors of both processes in three EHR-linked cohorts in the US (All of Us [AoU], n=599,423; Yale New Haven Health System [YNHHS], n=319,666; Michigan Genomics Initiative [MGI], n=82,372), reporting descriptive statistics for longitudinal visits and for a panel of 68 lab biomarkers commonly measured in EHRs. We conduct detailed model-based analyses of ten biomarkers spanning multiple domains: routine monitoring, general laboratory assessment, and symptom-triggered testing. These include glucose, hemoglobin A1c [HbA1c], creatinine, hemoglobin [Hgb], white blood cell count [WBC], low-density lipoprotein [LDL] and high-density lipoprotein [HDL] cholesterol, triglycerides, C-reactive protein [CRP], and thyroid-stimulating hormone [TSH]. Across the three cohorts, the median number of outpatient visits ranged from 1.7 to 6.1 per year over a median follow-up of 4.4 to 7.2 years. Among patients with at least one recorded measurement, the median within-person proportion of visits containing a given biomarker ranged from 0.4% to 19.5%, demonstrating that more frequent visits did not necessarily translate into greater per-visit biomarker capture. In the visit-process models, chronic disease burden, and a recent history of outpatient visits were consistently associated with higher visit rates across all three cohorts whereas associations with race, ethnicity, and neighborhood-level income varied across cohorts. In per-visit observation models, the association of covariates depended on the biomarker under consideration; for example, prior cancer diagnosis was associated with more frequent measurement of blood counts but with less frequent measurement of lipids. These findings provide a deeper understanding of how to model who seeks care and what is measured as two distinct recording processes in EHR. Our empirical findings show that the descriptors of these processes vary across cohorts and biomarkers, providing guidance on how to construct these models for downstream longitudinal analyses with irregular EHR visits.

15
Bridging the "Ten Walls" of Japanese Healthcare Data: A Comprehensive Semantic Mapping of JIPAD to HL7 FHIR R4 and Institutional Gap Analysis for the Japanese Health Data Space (JHDS)

Ohno, K.; Hashimoto, S.

2026-08-10 health informatics 10.64898/2026.08.06.26359847 medRxiv
Top 0.3%
6.2%
Show abstract

Background: Japan faces critical challenges in medical data interoperability, conceptualized as the "Ten Walls" obstructing the Japanese Health Data Space (JHDS) [1]. The Japanese Intensive Care Patient Database (JIPAD) - Japan's largest national ICU registry with 151 participating facilities - represents a high-quality critical care dataset that remains isolated from international data ecosystems. Objective: To develop a formal mapping of all 122 JIPAD variables to HL7 FHIR R4, characterize the nature and magnitude of semantic gaps, and assess the feasibility of JIPAD integration into the JHDS. Methods: All 122 JIPAD variables (Data Dictionary v3.7.2; Linkage Items List 20231020) were evaluated using ISO 21564 [8]-based semantic equivalence scoring across three tiers: High (direct FHIR R4 Core mapping), Partial (mapping via JP-Core Implementation Guide extensions [3]), and Low/No Equivalence (structural institutional gap). Semantically identical multi-instance fields (e.g., secondary disease codes x5) were consolidated into single mapping entries, yielding 114 mapping entries. Pseudonymization architecture was characterized from primary documentation. Results: Of 114 mapping entries representing the 122 JIPAD variables, 97 (85.1%) achieved High Equivalence via LOINC/SNOMED CT, and 12 (10.5%) achieved Partial Equivalence via JP-Core extensions, value-set translation, or FHIR R4 Core extension mechanisms - yielding a combined technical feasibility of 95.6% (109/114). Only 5 entries (4.4%) were classified as Low/No Equivalence, all attributable to Japan's proprietary disease classification system (288 adult codes; 165 pediatric codes) embedded in the DPC reimbursement framework, plus one Japan-specific procedure (PMX endotoxin adsorption) absent from international terminology systems. Variable-level mapping details are provided in Supplementary Table S1. Critically, JIPAD employs pseudonymization with record-linkage capability, enabling 99% DPC data matching - demonstrating that technical and design-level barriers to FHIR integration have already been resolved. Conclusion: JIPAD is technically and architecturally ready for FHIR integration at a 95.6% level. The remaining 4.4% barrier is exclusively institutional - rooted in MHLW policy frameworks governing the DPC disease classification system [6] - rather than technical. FHIR integration would further unlock pharmacoepidemiological and social epidemiological research currently inaccessible due to data isolation. As the sole national ICU registry providing high-acuity anchor data unavailable in general health records, JIPAD integration is essential for a clinically meaningful JHDS by 2027.

16
EpiKG2DAG: a Framework for Automated DAG Construction from Biomedical Text

DU, J.; Deng, G.

2026-08-11 health informatics 10.64898/2026.08.09.26360023 medRxiv
Top 0.3%
6.2%
Show abstract

While Directed Acyclic Graphs (DAGs) are essential for causal inference, their construction often relies on expert heuristics, which bypasses systematic evidence synthesis and creates a critical "evidence retrieval gap" in causal modeling. This study introduces EpiKG2DAG, a framework that supports evidence-anchored candidate DAG generation by transforming unstructured biomedical abstracts into structured epidemiological associations. We utilized DeepSeek-V3 to extract exposure-outcome association triplets from 189,266 abstracts and employed SapBERT for semantic normalization against UMLS concepts. The resulting Epidemiological Knowledge Graph (EpiKG) enables the automated identification of candidate confounders, mediators, and colliders based on graph-theoretic motifs and literature-derived evidence. A case study on COVID-19 and AKI demonstrates that the framework uncovers non-obvious confounders, such as air pollution, while ensuring evidence traceability. This work contributes to the field by mitigating the knowledge-acquisition bottleneck and providing a transparent, reproducible foundation for evidence-based causal modeling.

17
Local retraining mitigates domain shift in sepsis prediction: Lessons from translating a neonatal model to mixed intensive care data

Champeaux, S. A.; Booth, J.; Brown, A.; Sebire, N. J.; Drobnjak, I.; Bowyer, S.

2026-08-21 health informatics 10.64898/2026.08.18.26360666 medRxiv
Top 0.3%
5.5%
Show abstract

Background: Machine learning models leveraging electronic health records (EHRs) can support earlier detection of sepsis in intensive care units (ICUs). However, their clinical utility depends on reproducibility across institutions and patient populations. Building on a published pipeline from the Children's Hospital of Philadelphia (CHOP), this study examines how a neonatal sepsis prediction framework performs and can be adapted to a range of intensive care environments, paediatric, cardiac, and neonatal, at Great Ormond Street Hospital (GOSH). Methods: We extracted de-identified ICU EHR data from GOSH and applied feature derivation, unit harmonisation, and temporal sampling to align with the CHOP dataset used by Masino et al. (2019). Seven classifiers were first evaluated using CHOP-trained weights to characterise cross-domain behaviour and then retrained on local data to assess recoverability and site-specific adaptation. Model discrimination was summarised by AUC and F1, and learning curves were used to explore sample efficiency and bias-variance dynamics. Results: Models achieved strong discrimination on the CHOP neonatal cohort but demonstrated reduced performance when transferred to the mixed GOSH ICU population, reflecting anticipated domain and population shift. Retraining on GOSH data restored discrimination (AUC range 0.69-0.86), with Gradient Boosting (AUC 0.86 vs AUC 0.87 at CHOP) and KNN (AUC 0.80 vs AUC 0.79 at CHOP) models performing comparably to their CHOP benchmarks. DeLong's test confirmed statistically significant gains across all classifiers (p < 0.001). Conclusion: ICU cohort and baseline demographic differences between CHOP and GOSH introduced domain shift that limited direct model transfer. Elements of the original preprocessing pipeline could not be reproduced, further constraining transportability. Yet, retraining on local data restored high discrimination, showing that the modelling framework remains robust when re-estimated in new settings. These results highlight local adaptation as a practical route to recover performance and support safe, generalisable deployment of clinical prediction models in mixed clinical environments.

18
TrialCode Agent: LLM-Assisted Clinical Code-Set Construction for Trial Emulation

Habibdoust, A.; Sajjad, A.; Hernandez, D.; Patel, K.; Song, X.

2026-08-23 health informatics 10.64898/2026.08.20.26360962 medRxiv
Top 0.3%
5.4%
Show abstract

Objective Translating free-text clinical trial criteria into computable code sets is a valuable standardization practice that is necessary for producing reproducible real-world evidence studies but requires standardized interpretation across multiple clinical vocabularies. Methods We developed TrialCode Agent, a hybrid-large language model (LLM)-terminology verification agent that generates, formats, verifies, and expands candidate codes from free-text clinical criteria. The system supports ICD-9-CM diagnoses and procedures, ICD-10-CM, ICD-10-PCS, LOINC, and RxNorm medication concepts. We compared Baseline, Hybrid biomedical retrieval-augmented generation (RAG), and terminology-guided Family expansion pipelines using Claude, GPT Qwen, and MedGemma on 40 criteria from 11 trial groups. Performance was evaluated against expert-built reference code sets using exact-code precision, recall, and F1. Results The optimal pipeline varied by model. Claude with Baseline achieved the highest performance (precision 0.755, recall 0.619, F1 0.680), followed by GPT-5.5 with Baseline (precision 0.569, recall 0.658, F1 0.610), Qwen with Hybrid biomedical RAG (precision 0.656, recall 0.470, F1 0.548), and MedGemma with Family expansion (precision 0.487, recall 0.316, F1 0.383). Hybrid biomedical RAG improved aggregate F1 only for Qwen but increased GPT-5.5 RxNorm F1 from 0.320 to 0.909. Macro-averaged results showed criterion-level gains despite lower micro-averaged aggregate performance. Family expansion increased recall across models but generally reduced precision. In staged verifier ablation, micro-F1 increased from 0.254 before verification to 0.505 after final verification and expansion. Existence/vocabulary checking removed 2,594 false-positive codes, and acceptance filtering removed 952 additional false-positive codes before controlled expansion. Conclusions Combining LLM-based clinical interpretation with deterministic terminology verification produces auditable, database-ready code sets, but retrieval and broad family expansion do not consistently improve exact-code performance. Retrieval was particularly useful for RxNorm mapping, whereas overly broad or incomplete candidate generation remained the main source of error. Deterministic verification improves code validity and query readiness but cannot replace accurate clinical interpretation.

19
Reducing Under-Triage Risk in Large Language Model Based Clinical Triage Using UMLS-CUI Augmentation

Gokhale, R.; Kukreja, M.; Kumar, N.; Gourab, K.

2026-08-10 health informatics 10.64898/2026.08.07.26358932 medRxiv
Top 0.4%
5.3%
Show abstract

Background: Public facing large language models (LLMs) are increasingly used for health guidance, including triage recommendations. We evaluated whether augmenting LLM prompts with standardized clinical concepts from the Unified Medical Language System (UMLS) could improve the safety and robustness of clinical triage recommendations. Methods: We used a publicly available dataset comprising 60 clinician-authored clinical vignettes, each represented in 16 demographic and narrative variations, yielding 960 vignette-factor combinations. Clinical entities were extracted using a two-stage pipeline combining ClinicalBERT-based named entity recognition with rule-based identification of laboratory abnormalities. Extracted entities were mapped to UMLS Concept Unique Identifiers (CUIs). Negated concepts were excluded. A confidence-weighted CUI voting classifier was trained using empirical associations between CUIs and clinician-assigned triage categories. We compared five approaches: CUI-only classification, MedGemma 27B, MedGemma 27B augmented with CUIs, GPT-4o-mini, and GPT-4o-mini augmented with CUIs. Outcomes included overall accuracy, under-triage, over-triage, emergency-case accuracy, and sensitivity to anchoring statements. Results: CUI augmentation decreased under-triage but increased over-triage in both models tested (GPT-4o-mini and MedGemma 27B). It improved high-acuity recognition while reducing recognition of low-acuity cases. CUI augmentation had mixed effects on overall triage accuracy; accuracy increased for MedGemma 27B but decreased for GPT-4o-mini. Emergency-case accuracy improved from 73.0% to 80.7% for GPT-4o-mini and from 60.5% to 68.5% for MedGemma 27B. CUI augmentation also reduced susceptibility to anchoring statements. These findings suggest that the principal value of CUI augmentation may be shifting model behavior toward safety-oriented behavior rather than uniformly improving overall accuracy. Conclusion: Ontology-grounded prompt augmentation shifted LLM triage recommendations toward greater sensitivity to high-acuity presentations and reduced overall under-triage. These safety gains were accompanied by increased over-triage and mixed effects on overall accuracy. A hybrid architecture combining LLM-based language understanding with interpretable UMLS-derived clinical concepts may improve the safety and robustness of AI-assisted triage. Further evaluation using real-world patient communications and clinical outcomes is warranted.

20
Twelve-Year Real-World Evaluation of a Regulated Guideline-Based Warfarin Dosing and Care Automation System

Tiihonen, M.

2026-08-12 health informatics 10.64898/2026.08.10.26360059 medRxiv
Top 0.4%
5.2%
Show abstract

Background: Warfarin therapy requires repetitive dose adjustments based on INR (International Normalised Ratio) monitoring. We evaluated the long-term real-world performance of Forsante Warfarin Advisor (WA), a CE-marked class IIb guideline-based decision support and care automation medical device used in anticoagulation management. Methods: Retrospective real-world data from routine clinical use between 2016 and 2026 were analysed. Treatment quality was assessed using Time in Therapeutic Range (TTR). Recommendation performance was evaluated by comparing achievement of target INR after clinician acceptance or modification of Warfarin Advisor recommendations. Results: Among 1348 patients in March 2026 median TTR was 83%, compared with 70% in March 2016. Dosages congruent with Warfarin Advisor recommendations were strongly associated with achieving target INR at follow-up in INR target ranges of 2.0-3.0 and 2.5-3.5. Treatment quality remained consistently high across years of deployment. No serious device-attributable safety incidents, regulatory incident reports, or CAPA cases were identified during 12 calendar years and 82,709 patient years of routine use. Conclusions: The findings provide real-world long-term evidence that a guideline-based warfarin dosing and care automation system can support sustained high-quality anticoagulation control in routine clinical practice. The findings support the feasibility of deploying workflow-integrated execution of selected guideline-driven clinical processes, while the causal effects on clinical outcomes require prospective confirmation. Keywords: Clinical decision support systems, Guideline execution, Real-world evidence, Warfarin, Anticoagulation